Skip to main content

Client Registry / Master Patient Index

The client registry answers one question: are these two records about the same person? Everything else in a health information exchange depends on the answer being right.

The terms differ slightly by tradition — MPI (master patient index) is the hospital-sector term, client registry the OpenHIE term, EMPI (enterprise MPI) the term when it spans organisations — but the machinery is the same.


What it holds​

A small, deliberately constrained record per person:

  • Identifiers, each with its issuing namespace (national ID, health ID, facility MRNs)
  • Name, in the forms the culture actually uses
  • Date of birth, with a flag for estimated dates
  • Sex, and separately gender where relevant
  • Contact — phone, address, and their history
  • Mother's name, or another culturally appropriate discriminator
  • Links to source system records, and the confidence of each link
  • Merge history

It holds no clinical data. A client registry that starts storing diagnoses has become a shared health record and inherits all of that component's governance requirements.


Matching​

Deterministic matching​

Exact agreement on a defined set of fields, usually as an ordered rule set:

Rule 1: national_id matches exactly → MATCH
Rule 2: health_id matches exactly → MATCH
Rule 3: surname + given name + date_of_birth + sex all match → MATCH
Otherwise → NO MATCH

Fast, explainable, auditable. A clerk can be told why two records matched, and a court can be shown the rule.

Its weakness is that it fails on exactly the data that health systems have: transliterated names, estimated birth dates, single-name populations, phonetic spelling variation, and identifiers that are absent or wrong.

Probabilistic matching​

Each field contributes a weight based on how much agreement or disagreement on that field shifts the odds. Weights are summed into a score.

The insight from Fellegi–Sunter record linkage theory: a field's evidential value depends on how discriminating it is. Agreeing on an uncommon surname is strong evidence; agreeing on a common one is weak. Agreeing on sex is worth almost nothing, because half the population agrees by chance.

score
│
│ no match review match
├──────────────┬──────────────┬───────────────────▶
0 lower upper
threshold threshold
  • Above the upper threshold: automatic link
  • Below the lower threshold: automatic non-link
  • Between: human review

Techniques: string similarity (Jaro-Winkler, Levenshtein), phonetic encoding (Soundex, Double Metaphone, and language-appropriate equivalents — note that Soundex is built for English and performs badly on many languages), date tolerance, nickname and transliteration dictionaries, blocking to avoid comparing every record with every other.

Choosing​

DeterministicProbabilistic
ExplainabilityHighModerate — needs score breakdowns
Performance on clean dataExcellentExcellent
Performance on messy dataPoorBetter
Tuning effortLowOngoing
Governance burdenLowRequires a review team and threshold policy

Start deterministic if data quality does not support anything else, and record that as the reason. Probabilistic matching over data with 40% missing birth dates does not produce better matches; it produces confident wrong ones. Revisit the decision when the data improves — that is what the ADR is for.

Most mature registries are hybrid: deterministic rules on strong identifiers, probabilistic scoring for the rest.


Thresholds are a clinical safety decision​

Two error types, with asymmetric consequences:

ErrorConsequence
False positive — two people mergedOne person's allergies, results and medications appear in another's record. Direct patient-harm risk.
False negative — one person splitFragmented record; duplicate testing; missed history; inflated patient counts.

False positives are worse. Set thresholds conservatively, route uncertainty to human review, and make the review queue a funded role rather than a background task. The single best predictor of whether an MPI programme succeeds is whether anyone actually works the review queue.

Measure and publish: auto-match rate, review queue depth and age, duplicate rate in source systems, and — periodically, against a manually adjudicated sample — false positive and false negative rates.


The golden record​

The registry's consolidated view of a person, assembled from source records.

Survivorship rules decide which value wins when sources disagree: most recent, most trusted source, most complete, or human-adjudicated. Record the rule per field, and retain the source values — the golden record is a view, not a replacement. When a merge turns out to be wrong, the source values are what make recovery possible.


Merge and unmerge​

Record A ──┐
├──▶ Golden record (survivor) A and B remain, marked as
Record B ──┘ replaced-by, with full history

Requirements:

  • The losing identifier must continue to resolve — old references in old systems will keep arriving for years. FHIR handles this with Patient.link and the replaced-by type.
  • Downstream systems must be notified, and must be able to act on it. This is the hardest part: the shared health record, the HMIS, the insurer and every EMR each hold data keyed to the old identifier.
  • Unmerge must be possible. Merges are sometimes wrong, and discovering that after six months of clinical data has accumulated under the merged identity is a genuine emergency. If your platform cannot unmerge, that is a decisive selection criterion.
  • Every merge and unmerge is audited with the evidence and the deciding person.

The FHIR interface​

POST /Patient/$match # find candidates for a demographic set
GET /Patient?identifier=… # look up by identifier
GET /Patient/123/$everything # (on the SHR, not the registry)

$match is the standard operation and returns candidates with a match grade (certain, probable, possible, certainly-not) and a score. IHE PIX and PDQ profiles provide the equivalent in the pre-FHIR world and are still widespread.


Where it sits​

Registration desk ─────┐
CHW app ───────────────┼──▶ Interoperability layer ──▶ Client registry
Laboratory ────────────┘ │
┌──────────┴──────────┐
▼ ▼
Golden record Review queue
│ (humans)
▼
shared identifier returned to callers

The registry resolves identity; it does not store the encounter. See interoperability layer.


Practical guidance​

  1. Fix data capture before tuning the algorithm. A registration form that accepts a free-text birth date generates more duplicates than any matcher can repair. Enforce format, use pickers, add check digits, and show the clerk likely existing matches before they create a new record.
  2. Prevent duplicates at the point of creation. A search-before-create step is worth more than any downstream reconciliation.
  3. Load real data early. Matching performance is a property of your population's naming and data-quality patterns, not of the software.
  4. Test with adversarial cases: twins, common names, single-name populations, transliteration variants, name changes at marriage, estimated birth dates, and family members sharing a phone number.
  5. Publish the duplicate rate and hold source systems accountable for it.
  6. Plan the review team before go-live, including who covers it.

Open-source options​

ProjectNotes
OpenCRThe OpenHIE community client registry; FHIR-based, configurable deterministic and probabilistic rules
SanteMPI / SanteDBFull-featured MPI and health data platform, IHE and FHIR interfaces
OpenEMPILong-established; verify current maintenance before adopting
SplinkNot health-specific: an open-source probabilistic linkage library, useful for evaluating matching strategies against your own data before committing

All Tier 2. See platforms.


References​